Papers by C V Jawahar

4 papers
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency.
Approach: They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages.
Outcome: The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health).
IndicSpeech: Text-to-Speech Corpus for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: India has 22 languages, each of them being spoken by over a million people . the current state of the art text-to-speech systems for Indian languages are lacking in the multimedia domain .
Approach: They propose to train a state-of-the-art TTS system for Hindi, Malayalam and Bengali and publish the results.
Outcome: The proposed system trains neural text-to-speech systems for Hindi, Malayalam and Bengali and makes them publicly available.
More Parameters? No Thanks! (2021.findings-acl)

Copied to clipboard

Challenge: Using network pruning, we find that there are large redundancies in MNMT models.
Approach: They propose a method to prune and retrain redundant parameters of an MNMT model to improve bilingual representations while retaining multilinguality.
Outcome: The proposed method improves bilingual representations while retaining multilinguality.
CVIT’s submissions to WAT-2019 (D19-52)

Copied to clipboard

Challenge: In this paper, we explore multiway-models for Indian languages.
Approach: They propose to use a Transformer architecture to experiment with multilingual models and methods for low-resource languages.
Outcome: The proposed system is feasible in low-resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations